National Science Review
◐ Oxford University Press (OUP)
Preprints posted in the last 90 days, ranked by how well they match National Science Review's content profile, based on 21 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
Li, D. J.
Show abstract
All cellular life forms fall under the three-domain classification of life, raising a fundamental evolutionary question: why does this classification feature three rather than two or four? To answer this question, a more general method, rather than the traditional one based on comparing small-subunit ribosomal RNAs, is required. The three-base periodicity in genomes is a common feature of both cellular life forms and viruses, which is species-specifically biased between amino acid biosynthetic families. Based on comparing such a common feature of all life forms, a global triangular diversification picture has been obtained, whose three angular regions correspond to the three domains, respectively. This mechanism of diversification of life attributes the evolutionary driving forces in diversification of the three domains of life to the biases between amino acid biosynthetic families. Notably, the same mechanism also applies to the contemporary diversification of SARS-CoV-2, whose reasonable results in turn corroborate the above explanation of primordial diversification of life and in addition shed light on the mechanism of speciation.
Sun, Y.; Yao, W.; Zhang, J.; Song, W.; Zhao, X.; Hao, C.; Chen, X.; Zeng, S.; Jia, S.; Yang, Y.; Chen, X.; Xiao, X.; Poo, M.-m.; Sun, Y.; Xu, B.; Zhang, T.
Show abstract
The organizational principles of natural neural networks could inspire the new architecture design of artificial neural networks (ANNs). Analysis of single-neuron connectomes of mouse brains revealed distinct profiles of three-node connectivity motifs in various cortical areas and hippocampal formation. A connectome-informed neural network algorithm ("CINA") was developed to incorporate natural connectivity motifs into ANN algorithms represented by recurrent neural network (RNN) and transformer-based large language model (LLM). We found that incorporation of the average profile of cortical motifs improved the RNNs performance in noise-resistant categorization and motor learning benchmark tasks, as compared with RNNs with random connectivity. Notably, incorporating cortex-specific motifs further elevated the RNNs performance in tasks related to the cortical function, and this effect was enhanced by artificially increasing the bias in the motif profile. Similar experimental results were verified on an LLM using Motif-Transformer for natural language question answering and brain-signal decoding tasks. Graph-theoretic analyses showed that incorporating natural motifs drove the emergence of modular and small-world properties in ANNs. Together, we demonstrated not only connectome-inspired optimization of ANN architecture but also functional significance of specific motif profiles in various cortices.
Liao, H.; Qin, B.; Zhou, L.
Show abstract
Objectives; The role of nuclear receptor subfamily 4, group A, member 3 (NR4A3) in hepatic steatosis, inflammation, and insulin resistance (IR) within the context of metabolic dysfunction-associated steatotic liver disease (MASLD) remains largely underexplored. Consequently, this study aimed to examine NR4A3's impact on MASLD and the potential underlying mechanisms. Methods; We aimed to elucidate the functional role of NR4A3 in MASLD through its knockdown in cell culture and animal models. To establish the cell culture model of MASLD, LO2 cells were treated with free fatty acids (FFAs), while male C57BL/6 mice were fed a high-fat diet (HFD) to create the animal model. NR4A3 knockdown was achieved using specific short hairpin RNA (NR4A3-shRNA) in the mice model and three small interfering RNAs (NR4A3-siRNAs) in the cell culture model. The lipids content, fatty acid synthesis, inflammatory factors, and IR were then assessed with and without NR4A3 knockdown. Furthermore, the underlying mechanism through which NR4A3 exerts its influence was explored by analyzing the interaction between NR4A3 and activating transcription factor 3 (ATF3). Results: In the cell culture experiments, the knockdown of NR4A3 significantly decreased the lipids content, fatty acid synthesis, and inflammatory factors in the LO2 cells treated with FFAs in the NR4A3-shRNA group compared with those in the NC-shRNA control group. In the animal model experiments, NR4A3 knockdown in the HFD male C57BL/6 mice significantly ameliorated HFD-induced hepatic steatosis, inflammation, and IR. Mechanistically, the knockdown of NR4A3 downregulated the expression and transcriptional activity of ATF3, resulting in an impaired ATF3 function. ATF3 overexpression significantly reversed lipid accumulation decline and reduced inflammation after NR4A3 knockdown. Conclusion: The downregulation of NR4A3 alleviates MASLD by modulating ATF3, suggesting this may be a promising therapeutic target.
Zhao, D.; Yang, Y.; Sun, J.; Zhang, J.; Duan, H.; Tan, Y.; Liu, l.
Show abstract
Although the "RNA world" hypothesis suggests that RNA played a crucial role in the origin of life [7], the functional framework of RNA in prebiotic protein synthesis and the mechanisms of genetic code formation during the prebiotic period remain poorly understood. Here, using the prebiotic "primordial soup" as a model, we reconstructed the detailed steps that would yield a protein with a stable ordered amino-acid sequence in the "primordial soup" at the prebiotic period. In the "primordial soup", a large number of medium- to large-sized biomolecule-like substances--such as RNA-like and protein-like molecules of various sizes and shapes, as well as related polymers like amino-acid-RNA-like etc.--did generate and accumulate. Moreover, protein-like and RNA-like molecules formed even more intricate complexes. These complexes bound free mRNA-like molecules through complementary base pairing. Subsequently, with an extremely low probability, two adjacent amino-acid-RNA-like molecules became bound to this free mRNA-like molecule, and their amino acids underwent a condensation reaction by the complexes, producing peptides and eventually proteins or polypeptides. This free mRNA-like molecule exhibits a certain flexible structure, whereas the super-large complexes formed by protein-like and RNA-like molecules (which possess certain activities) and the amino-acid-RNA molecules exhibit relatively rigid structures. Long-term evolution and mutual selection led to the emergence of proteins with stable amino acid sequences and moderate catalytic activity. In this way, the nucleotide information embedded in such mRNA-like molecules indirectly express through protein synthesis--a process we term the "A Co-Adaptation Flexible-Rigid Docking Model", where flexible mRNA-like molecules dock onto rigid complexes to enable ordered peptide formation. Finally, we show how trinucleotide codons emerge naturally from the flexible-rigid docking constraints.
Niu, Q.; Su, M.; Liang, L.; Che, Z.; Zhu, Q.; Wang, F.; Xiao, J.
Show abstract
Background Alcohol-associated liver disease (ALD) has emerged as a major cause of chronic liver disease and liver-related mortality in China. This study aimed to project the future burden of ALD in Chinese adults from 2020 to 2050, including prevalence of ALD, number of alcoholic steatohepatitis (ASH) cases, incident hepatocellular carcinoma (HCC) cases, liver transplantation (LT) demand, liver-related deaths, and disability-adjusted life years (DALYs). Methods We developed an agent-based state-transition microsimulation model with yearly cycles and a lifetime horizon. The model simulated 5,678,912 representative Chinese adults (mean age 36.2 years, 51.2% male). Health states included no steatosis, alcohol-associated steatotic liver, ASH, fibrosis stages F0-F4, decompensated cirrhosis, HCC, LT, and liver-related death. Model inputs were derived from the China Kadoorie Biobank, Global Burden of Disease Study 2021, China's national surveys, published meta-analyses, and transplant registry data. Projections incorporated demographic shifts, alcohol consumption trends, and calibrated transition probabilities. Uncertainty was assessed via 1,000 Monte Carlo simulations generating 95% uncertainty intervals. Results ALD prevalence was projected to increase from 4.8% (55 million individuals) in 2020 to 8.5% (94 million individuals) by 2050. ASH cases rose from approximately 18 million to 20 million. Annual incident HCC cases nearly doubled from 20,500 in 2020-2025 to 45,200 by 2046-2050. LT demand quadrupled from 2,300 to 9,800 cases. Liver-related deaths increased from 50,000 in 2020 to 85,000 in 2050, while DALYs rose from 1.5 million to 2.6 million. Conclusions In the absence of strengthened alcohol control policies, ALD will impose a substantial and growing burden on China's health system by 2050, with marked increases in HCC incidence, LT demand, and liver-related mortality.
Zuo, N.; Cai, X.; Wang, W.; Ren, Z.; Jiang, Z.; Jiang, W.; Song, X.; Gu, Y.
Show abstract
Nicotine accumulates in the gut and drives non-alcoholic steatohepatitis (NASH) via the gut-liver axis, yet no effective clinical intervention is currently available. To address this challenge, the probiotic Escherichia coli Nissle 1917 (EcN) was engineered for in situ nicotine clearance in the gut. Mutational screening of nicotine oxidoreductase 2 (PpNicA2) identified a highly active variant, PpNicA2A107R. Its incorporation into EcN together with an electron transfer protein (CycN) and a newly identified transporter (T3/T7) yielded 80% nicotine-degrading activity. Chromosomal integration of this module generated a stable strain, EcN-N12, which in NASH mouse models depleted intestinal nicotine, rescued hepatic lipid metabolism, alleviated tissue damage, and intercepted the nicotine-mediated gut-liver axis pathological progression. This work thus offers an effective and clinically translatable approach for nicotine-associated diseases.
Chen, F.
Show abstract
Viral infection induces tRNA competition between viral and host genes, often repressing host translation. However, how endogenous genes are affected by this competition remains unclear. Three possible hypotheses are considered: abundant-tRNA shortage, rare-tRNA shortage, and viral similarity repression. Pan-virus Ribo-seq data show that endogenous genes with codon usage bias (CUB) matching host tRNA supply or viral CUB are strongly repressed, due to a positive correlation between endogenous CUB-tRNA mismatch and endogenous-viral CUB difference, supporting the abundant-tRNA shortage and viral similarity repression hypotheses. In E. coli experiments with synonymous gentamicin resistance proteins, this positive correlation supports abundant-tRNA shortage, while a non-positive correlation supports rare-tRNA shortage, and both positive and non-positive correlation types support viral similarity repression. Finally, analysis of human virus genomes reveals this positive correlation for most viruses, but a non-positive correlation in a few, reflecting diverse virus-host interaction strategies. These findings establish viral similarity repression as a universal principle, uncovering previously unrecognized complexity in virus-host coevolution.
Zhang, Z.; Feng, Y.; Ge, X.; Meng, X.; Peng, Y.
Show abstract
Viral phenotypes such as host and tissue tropism are critical determinants of viral infection and transmission. Inferring viral phenotypes presents unique challenges compared to cellular organisms, as viruses rely entirely on host machinery for replication and survival. Current methods for predicting viral phenotypes mainly rely on viral genomic data, often overlooking host-related information. Here, we evaluated the utility of predicted virus-human protein-protein interactions (PPIs) in inferring diverse viral phenotypes using machine-learning algorithms. For predicting human infectivity, a PPI-based machine learning model outperformed both virus genomic and protein sequence-based models that used large language model embeddings. It also surpassed previous methods that incorporated both viral and host genomic data. The human proteins identified by the model were significantly enriched in functions related to viral infection and immune response. In predicting various phenotypes of human RNA viruses, PPI-based models performed better than virus sequence-based models in forecasting virulence, human transmissibility and transmission routes, while showing comparable performance to genomic sequence-based models in predicting tissue tropism. Finally, we demonstrated that a PPI-based model could distinguish high-risk HPV genotypes from low-risk ones. Proteins associated with high-risk HPV were involved in apoptosis and immune regulation, whereas those linked to low-risk HPV were enriched in telomere maintenance and DNA repair. Collectively, this study is the first to demonstrate the value of predicted virus-human PPIs in inferring viral phenotypes, thereby enhancing our understanding of the molecular mechanisms underlying these phenotypes. It also provides effective tools for risk assessment of emerging viruses, contributing to improved pandemic preparedness.
Zhang, X.; Zhang, F.; Suonan, Z.; Zhang, Y.; Li, Y.-L.; Li, X.; Kim, S. H.; Zhou, Y.; Lee, K.-S.; Yu, L.
Show abstract
While life-history strategies are typically fixed within species, evolutionary transitions between perenniality and annuality can occur. In clonal seagrasses, annual and perennial plants often coexist in the same population, providing a unique model for studying the genetic basis of this transition. Two seagrass Zostera marina populations in South Korea display a striking dichotomy: shallow-water sub-populations follow a typical perennial strategy, whereas their deep-water counterparts are annual. Here we show that this shift from perenniality to annuality, potentially caused by the SAPK7 gene, is genetically coupled with the CAO gene, which is under strong positive selection for low-light adaptation. The up-regulation of the SAPK7 gene triggers early flowering in seedlings, before the formation of any lateral shoots via asexual reproduction. In this special case where the genet contains only one ramet, the post-reproductive death of the ramet is equivalent to the death of the whole genet, which explains the annual phenotype. Our findings reveal a mechanistic example where annuality overcomes perenniality by hitchhiking on a positively selected gene. Given that the ancestral state of plants is perennial, this coupling of annuality with beneficial alleles may represent one of the pathways for the repeated evolution of annual life histories across flowering plants.
Zhang, T.; Xiong, Y.; Chen, K.; Wu, S.; Yan, X.; Zhou, J.; Wang, Y.; Yang, C.; Wang, P.; Zhou, Z.
Show abstract
Camptothecin derivatives are first-line anticancer drugs used worldwide for the treatment of diverse malignant tumors. However, the biosynthetic pathway of camptothecin has remained elusive for five decades. Here, we fully map its entire biosynthetic route. We discovered five key missing enzymes (OpCAR, OpSDR11, OpCS, OpGH1, and OpSTR) via the combination of MALDI mass spectrometry imaging, single-cell RNA sequencing and co-expression analysis. Meanwhile, we demonstrated a free flavin mononucleotide triggered the non-enzymatic 6-5-6 to 6-6-5 fused-ring skeleton rearrangement, filling the last gap in camptothecin biosynthesis. Finally, we validated this identified pathway and achieved the de novo biosynthesis of camptothecin in Saccharomyces cerevisiae. These discoveries uncover the long-standing mystery underlying camptothecin and pave the way for manufacturing camptothecin and its derivatives through synthetic biology approaches.
Zhang, S.;Tyshkovskiy, A.;Ying, K.;Wang, S.;Gladyshev, V.
Show abstract
Alternative splicing exhibits significant changes during development and aging, affecting the composition and variance in the transcriptome. However, it is unclear whether and how age-associated splicing dysregulation leads to functional consequences. Here, an integrative analysis of transcriptome data across mouse and human tissues revealed that aging is characterized by systematic deterioration of the fidelity of RNA splicing, here termed splicing degeneration, a measure of functional alteration of reading frame and domain configuration of protein products. Genes with higher aging-associated splicing degeneration were more conserved and enriched for processes such as RNA metabolism and antigen presentation. By assessing alternative splicing events associated with functional deterioration, we quantified the degree of splicing degeneration. Its level increased with age but was alleviated following calorie restriction or rapamycin treatment, indicating that it can serve as a new molecular hallmark of aging. Mechanistically, through a comprehensive meta-data analysis, we discovered that splicing degeneration is associated with age-associated changes in specific splicing factors, which in turn showed a strong association with age-related transcriptome changes. Overall, our study demonstrates the intricate relationship between aging and genome-wide splicing degeneration, revealing a promising target for aging interventions acting to reverse splicing degeneration.
Wang, X.; Gao, Q.; Ge, L.; Qin, Z.
Show abstract
BackgroundThe biosynthesis of bacterial aromatic polyketide polyketides (type II polyketides, T2PKs) employs a single set of catalysts (ketosynthases, KSs or KS, with chain length factors, CLFs or KS{beta}) and iteratively assembles a carbon backbone with precise chain length control. Considering the increasing number of T2PKs discovered in laboratory settings, it is necessary to understand the evolution trajectories of KSs and CLFs. ResultsWe employed our recently developed algorithm, MAAPE, based on large protein language model (PLM) to glean insights into the evolution process of KSs and CLFs. Our findings indicated the evolutionary history of KS and CLF domains from bacterial T2PKSs and identified a shared ancestral cluster (Cluster A), supporting a common origin. Despite structural homology, KSs and CLFs followed distinct evolutionary paths, shaped by coevolution and early horizontal gene transfer. ConclusionsUnderstanding the evolutionary lineage of these enzymes will illuminate the natural optimization processes of their functions and present opportunities for the rational design of novel polyketides with enhanced efficacy.
Huo, Y.; Li, J.; Huang, J.; Dong, Z.
Show abstract
Abstract Objective Commercially available next-generation sequencing (NGS) platforms in China routinely adopt a lung cancer-derived tumor microenvironment (TME) subtyping signature from a European cohort to classify colorectal cancer (CRC), yet its diagnostic performance in Chinese CRC patients remains unvalidated. This study aimed to evaluate the subtyping efficiency of the lung cancer TME signature in a Chinese CRC cohort, screen CRC-specific immune mRNA biomarkers for TME subtyping, and explore the clinical utility of IRF1, CD8A and LAG3 for distinguishing immune-enriched (IE) and immune-desert plus fibrotic (D+F) subtypes. Methods A total of 87 FFPE CRC specimens with complete NGS and clinicopathological data were retrospectively enrolled, including 15 IE subtype and 72 D+F subtype patients. Thirty-one mRNA transcripts covering 13 immune-metabolic homeostasis genes and 18 immune checkpoint/infiltration-related genes were divided into two functional modules. Spearman correlation analysis was performed to assess co-expression patterns among candidate genes. Receiver operating characteristic (ROC) curves combined with five-fold cross-validation were used to compare the discriminatory efficacy of single-gene markers and the three-gene combined panel. Results Strong positive co-expression was observed between IRF1, CD8A and LAG3 (IRF1-CD8A: r=0.93; IRF1-LAG3: r=0.84; CD8A-LAG3: r=0.73). Nominal P-values indicated elevated expression of IRF1, CD8A and LAG3 in IE subtype, though no intergroup significance remained after Benjamini-Hochberg FDR correction, largely attributed to the limited sample size of IE cases. Single-gene ROC analysis showed AUC values of 0.763 (IRF1), 0.752 (CD8A) and 0.771 (LAG3), with LAG3 exhibiting the best individual discriminatory capacity. The three-gene combined panel yielded a cross-validated AUC of 0.717, inferior to single LAG3, due to severe collinearity that generated redundant predictive information. Conclusions The lung cancer-originated TME subtyping system cannot be directly extrapolated to Chinese CRC patients. LAG3 serves as a promising independent transcriptomic candidate marker for distinguishing CRC TME subtypes. The robust collinearity among IRF1, CD8A and LAG3 eliminates additional predictive benefits of the combined signature. Large independent multi-center Chinese CRC cohorts are required to construct population-specific immune transcriptomic biomarkers for standardized clinical NGS TME stratification.
Tian, X.; Fung, A. A.; Shang, X.; Zhang, D.; Chen, B.; Zhang, L.; Li, K.; Zhong, M.; Deng, Y.; Yang, M.; Lu, Y.; Tao, B.; Gao, F.; Baysoy, A.; Lin, X. L.; Ivovic, A.; Chen, S.; Li, F.; Xu, M. L.; Zhang, X.; Gerstein, M.; Yang, X.; Liu, C.; Fan, R.
Show abstract
Glycosylation is a fundamental process regulating cellular function, tissue organization, and disease progression. However, comprehensive glycan profiling at single-cell spatial resolution remains largely inaccessible, particularly in clinical archival tissues. Here we develop spatial-GPT, a multimodal platform for simultaneous profiling of glycans, proteins, and/or transcripts in archival formalin-fixed paraffin-embedded (FFPE) tissues. Using a panel of 30 DNA-encoded lectins recognizing major mammalian glycan motifs and structural classes, sequencing-based spatial-GPT (DBiT-GPT) mapped the spatial glycome, proteome, and transcriptome across 16 human liver specimens encompassing steatosis, fibrosis, cirrhosis, and hepatocellular carcinoma (HCC), leading to identification of spatial glyco-codes - combinatorial glycan states associated with distinct cellular identities, tissue features, and pathological processes. Unexpectedly, glyco-codes alone were sufficient to resolve major cell types, disease states, and HCC subtypes, revealing a previously unappreciated level of biological information encoded within the tissue glycome. Spatial glycomics uncovered tumor-like glyco-codes in premalignant regions, suggesting that glycan reprogramming may precede overt malignant transformation. Using imaging-based single-cell spatial glycan-protein profiling (CODEX-GP), we track glyco-codes across the whole-tissue architecture of 3 representative HCC samples. We further examined the glyco-codes across more than 300 patient specimens and quantified cell-type- and disease-specific glyco-codes as well as glycan-defined immune-evasion, T-cell-exhaustion, and steato-fibrotic niches. Together, these findings establish spatial glyco-codes as a previously unrecognized layer of tissue organization that encodes cellular identity, tissue function, and disease progression. The ability of glyco-codes to distinguish major liver pathologies across independent patient cohorts further highlights their potential as a new class of molecular histopathology biomarkers.
Nasrolahpour, H.; Jandera, A.; Skovranek, T.; Despotovic, V.; Pellegrini, M.
Show abstract
Epigenetic clocks based on DNA methylation patterns are among the most accurate molecular correlates of chronological age, yet widely used clocks are predominantly empirical models with limited explicit characterization of the underlying methylation variability, lacking a direct connection to the physical mechanisms of aging. In this work, we bridge this gap by introducing an information-theoretic framework for DNA methylation dynamics combined with nonlinear machine learning to develop a competitive and interpretable age predictor. We model the population distribution of methylation {beta}-values at each CpG site using a reparameterized three-parameter Generalized Gamma Distribution (GGD) and derive a closed-form expression for its differential Shannon entropy. The resulting CpG-level entropy is used to characterize methylation variability and as a criterion for locus filtering. We introduce the Stacy Gradient Boosting Clock (Stacy-GB), which combines this GGD-based representation with a LightGBM regressor. The model was evaluated across independent cohorts using the ComputAgeBench epigenetic clock benchmark. Stacy-GB achieved a mean absolute error (MAE) of 3.74 years and a median error (bias) of 2.41 years, significantly outperforming state-of-the-art epigenetic clock baselines. Furthermore, age acceleration estimated by Stacy-GB was associated with several clinical pathologies, including ischemic heart disease, HIV infection, multiple sclerosis, and Werner syndrome, supporting its potential as an accurate and biophysically grounded tool for clinical aging research.
Xiao, W.; Zheng, Q.; Wang, Y.; Yuan, Y.; Chen, Y.; Zheng, T.; Chen, Y.; Gao, Y.; Song, B.; Zhang, B.; Qiu, L.; Zeng, L.; Huan, M.; Brown, C. H.; Duan, S.; Pan, G.; Gao, Z.
Show abstract
Oxytocin-mediated milk ejection (ME) is pivotal to effective breastfeeding and productive health, yet behaviorally decoding and revealing neural mechanisms of ME remains challenging. Here, we combined in vivo calcium imaging and intramammary pressure recording to uncover the temporal connections between episodic activity of oxytocin neurons and ME in conscious lactating rats. Leveraging the association and behavioral responses in dam and pup, we developed a supervised machine learning framework (ME Decoder) to enable automated analyses of ME. Inspired by its interpretable features, we defined the activity-coupled dam-pup interactions (ADPI), manifested by high kyphosis of the dam followed by pup treading and stretch, as the behavioral signatures of ME. By ME Decoder and ADPI analyses, we detected reduced ME but unaffected activity of oxytocinergic neurons after systemic blockade of oxytocin receptor. Our study uncovers the oxytocinergic and behavioral signatures of ME and provides a generalizable approach for further investigation.
Liu, T.; Zhai, S.; Lin, S.; Zhan, X.; Deng, J.; Liu, H.; Siu, S. W. I.
Show abstract
Harnessing enzyme specificity requires a thorough understanding of enzyme promiscuity, which determines enzymes catalytic scope; however, measuring this scope still relies heavily on labor-intensive analytical approaches. While data-driven approaches have emerged to predict the catalytic scope of enzymes, these methods continue to face challenges such as restricted datasets and insufficient integration of enzyme structural information and reaction transformations. Here, we introduce MAERM, an innovative mixed-attention model designed to predict enzyme-reaction matching relationships. Built on our MAERM-DB, a dataset with broad coverage of validated and chemoenzymatic catalysis data, MAERM utilizes a local-global attention module to integrate multimodal enzyme information with fine-grained reaction representations, thereby predicting enzyme-reaction matching probabilities. Results show that MAERM consistently outperforms all baselines, with an average F1-score of 0.984. Notably, on challenging test samples with less than 40% sequence identity to the training set, MAERM outperforms the second-ranked model by 5.9% in F1-score. In addition, MAERM achieves the highest top-10 success rate of 51.7% on Enzyme-405 and the highest balanced accuracy of 0.697 on BioCat-547, further supporting its generalizability in enzyme screening and chemoenzymatic catalysis. Finally, MAERM can serve as an efficient scoring module. When integrated with ProteinMPNN, MAERM has successfully guided novel enzyme design for two carbonyl reduction reactions, resulting in enhanced catalytic potential for the native substrate and demonstrating broad compatibility. Overall, MAERM has the potential to reduce the experimental cost of measuring enzymes catalytic scope, facilitate enzyme design, and ultimately accelerate the design-build-test-learn cycle in enzyme engineering.
Tsui, D.; Talreja, K.; Aghazadeh, A.
Show abstract
Understanding how mutations combine to shape protein fitness remains a central challenge in biology, driven in part by the prevalence of highorder epistasis. Existing analyses of epistasis, however, implicitly define epistatic interactions under a uniform probability measure over sequence space, even though evolution constrains natural proteins to a highly structured, non-uniform distribution of sequences. Here, we show that the apparent complexity of protein epistasis depends fundamentally on the underlying evolutionary distribution of sequences. We develop an evolution-aware spectral framework that incorporates the evolutionary distribution of amino acids at each sequence position, inducing an orthogonal decomposition under the evolutionary measure while preserving efficient spectral algorithms for scalable analysis. Across diverse protein fitness landscapes, this framework consistently produces more compact spectral representations, explaining more phenotypic variation with fewer epistatic interactions while substantially reducing apparent high-order epistasis. It also enables more accurate recovery of fitness landscapes from limited experimental measurements and concentrates the remaining higher-order interactions into localized, structurally interpretable motifs. These results suggest that a substantial fraction of apparent high-order epistasis arises from defining epistatic interactions under a uniform measure over sequence space and can be resolved by aligning spectral analysis with evolutionary constraints.
Liu, Y.
Show abstract
The Gene Version Iteration Hypothesis (GVIH) proposes that mutant genes may originate from the Y chromosome, traverse through the X chromosome to autosomes, undergo interchromosomal transfer, and potentially return to the Y chromosome via the X chromosome. This hypothetical closed transmission loop may facilitate the storage, screening, and elimination of different versions of mutant genes. The hypothesis comprises five core propositions: (1) Mutation reservoir: The Y chromosome may serve as a specialized carrier for generating mutant genes, characterized by elevated mutation rates, reduced gene density, and accelerated evolutionary dynamics; (2) Closed-loop transmission: Mutant genes may follow a unidirectional pathway Y[->]X[->]autosomes[->]X[->]Y, forming a complete transmission circuit; (3) Coexistence of multiple versions: A single functional gene may exist in multiple versions across different chromosomes, constituting a dynamic gene version library; (4) Reproductive screening: Environmentally adaptive gene versions may persist across generations and potentially migrate to upstream chromosomes, while maladaptive versions may be eliminated; (5) Terminal elimination: Gene versions reaching the Y chromosome may undergo elimination processes, potentially preventing version monopolization and maintaining evolutionary dynamics. This hypothesis provides a novel framework for understanding adaptive evolution at the genetic level. If empirically validated, it may offer new insights into the molecular mechanisms underlying certain genetic phenomena and evolutionary processes.
Price, T. A.; Liu, A.; Cowan, R. L.; Shahdoust, N.; Davis, T. S.; Kundu, B.; Rolston, J. D.; Rahimpour, S.; Shofty, B.; Borisyuk, A.; Smith, E. H.
Show abstract
Real-world decision-making rarely occurs with perfect information. Instead, individuals must constantly weigh potential rewards against the probability of adverse outcomes.1 Failures of this process can lead to maladaptive decisions associated with reduced lifetime success, and numerous psychiatric disorders such as gambling addictions, bulimia nervosa, and substance use disorder.2,3 The neural computations that facilitate inference about the landscape of potential outcomes remain unclear, but are thought to occur in distributed frontotemporal circuits.4 Here we used deep reinforcement learning agents to predict distinct behavioral strategies and their underlying neural population dynamics during a risky decision-making task. Across a range of training conditions, deep reinforcement learning agents separated into strategies marked by either overly cautious exploration of the reward contingency space or a high-performing, risk-adaptive Bimodal strategy. The internal dynamics of high-performing Bimodal agents formed low-dimensional representations that segregated safe and risky states. In contrast, the cautious exploration agents were associated with more skewed and entangled neural representations. We found remarkably similar dynamical representations and their associated behavioral strategies in neuronal ensemble recordings from human epilepsy patients performing a similar risky decision-making task. These results reveal the structure of dynamical computations that underlie inferences about uncertain outcomes and their associated behavioral strategies.